Papers with safety rate
Athena: Safe Autonomous Agents with Verbal Contrastive Learning (2024.emnlp-industry)
Copied to clipboard
| Challenge: | Existing safety benchmarks on the ability of large language models to perform tasks are lacking. |
| Approach: | They propose a framework that leverages verbal contrastive learning to guide agents towards safety . they use past safe and unsafe trajectories as in-context examples to guide them towards safety. |
| Outcome: | The proposed framework leverages verbal contrastive learning to guide agents towards safety while performing tasks. |
PerMemSafe: Benchmarking Implicit Personalized Safety of Long Horizon Self-Evolving Agents (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing self-evolving agents have a low safety rate in long-horizon interactions . however, this reliance on context-independent safety evaluations is insufficient . |
| Approach: | They propose a framework that explicitly models personalized risk inference and memory evolution. |
| Outcome: | The proposed framework improves implicit personalized safety by 23.8% over prior frameworks while maintaining helpfulness in long-horizon interactions. |